Goto

Collaborating Authors

 evaluation function


François Pachet on music generation with AI

AIHub

Dr François Pachet is an AI researcher and musician, and one of the most influential figures in AI and music. His innovative contributions have defined the field over the past decades through creative systems such as the Continuator, and Flow Machines, among others. After leading the Spotify Creator Technology Research Lab and the Sony Computer Science Lab, he went on to create his own companies: Imagine All The People and Ynosound. In the context of IJCAI2025, he spoke about what deep learning changed, and what still remains wide open. He explains why tools like Suno and Udio--ChatGPT-like platforms for music generation--can produce astonishing results that still feel unsatisfying; why the next step for music generation requires combining sampling with search; and why the most important problems in artistic domains are, by nature, ill-defined--because there is no loss function to determine what is "good". Above all, he defends the importance of researcher autonomy: work on the questions that genuinely fascinate you, even when they fall outside prevailing trends--perhaps especially then. Thank you for joining me for this interview. Could you begin by telling us when was your first IJCAI and a memory related to it? I think the first IJCAI I attended was in Montreal in '95. I was there for a couple of workshops, one of them was about music and AI, and the other one I think was on advisor systems, something like that. And I remember there was a French colleague who was there also, at the time he was doing his PhD. And there was this researcher called Herbert Simon, who is a Nobel Prize pioneer of AI. I remember having chatted a little bit with this French guy who was very bold, and he just went up to Simon, said "Hey," and he started a conversation with him. And I was very impressed by the fact that you could meet those kinds of guys informally in a corridor or something at this conference.


Influence Functions for Edge Edits in Non-Convex Graph Neural Networks

Neural Information Processing Systems

Understanding how individual edges influence the behavior of graph neural networks (GNNs) is essential for improving their interpretability and robustness. Graph influence functions have emerged as promising tools to efficiently estimate the effects of edge deletions without retraining. However, existing influence prediction methods rely on strict convexity assumptions, exclusively consider the influence of edge deletions while disregarding edge insertions, and fail to capture changes in message propagation caused by these modifications. In this work, we propose a proximal Bregman response function specifically tailored for GNNs, relaxing the convexity requirement and enabling accurate influence prediction for standard neural network architectures. Furthermore, our method explicitly accounts for message propagation effects and extends influence prediction to both edge deletions and insertions in a principled way. Experiments with real-world datasets demonstrate accurate influence predictions for different characteristics of GNNs. We further demonstrate that the influence function is versatile in applications such as graph rewiring and adversarial attacks.


Self-Challenging Language Model Agents

Neural Information Processing Systems

Large language models are quickly becoming the foundation for intelligent agents that are capable of using tools. However, training such agents is challenging because it requires human creation and annotation of a diverse set of tasks, tools, and evaluation criteria. In this paper, we propose the Self-Challenging framework for training an agent on high-quality tasks that are generated by itself. The agent first plays the role of challenger and generates a task after interacting with the given tools. The tasks take the form of a novel general class of problems termed Code-as-Task, which are defined by an instruction, a verification function and solution and failure cases which serve as tests, allowing to filter only for highquality tasks. The agent then takes an executor role and trains on those tasks with reinforcement learning using the evaluation feedback as a reward. Evaluation on two existing multi-turn tool-use agent benchmarks, M3ToolEval and TauBench, shows the Self-Challenging framework achieves over a two-fold improvement in Llama-3.1-8B-Instruct,


Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models

Neural Information Processing Systems

Text-to-image diffusion models rely on text embeddings from a pre-trained text encoder, but these embeddings remain fixed across all diffusion timesteps, limiting their adaptability to the generative process. We propose Diffusion Adaptive Text Embedding (DATE), which dynamically updates text embeddings at each diffusion timestep based on intermediate perturbed data. We formulate an optimization problem and derive an update rule that refines the text embeddings at each sampling step to improve alignment and preference between the mean predicted image and the text. This allows DATE to dynamically adapts the text conditions to the reverse-diffused images throughout diffusion sampling without requiring additional model training. Through theoretical analysis and empirical results, we show that DATE maintains the generative capability of the model while providing superior text-image alignment over fixed text embeddings across various tasks, including multi-concept generation and text-guided image editing.




Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models

arXiv.org Artificial Intelligence

Text-to-image diffusion models rely on text embeddings from a pre-trained text encoder, but these embeddings remain fixed across all diffusion timesteps, limiting their adaptability to the generative process. We propose Diffusion Adaptive Text Embedding (DATE), which dynamically updates text embeddings at each diffusion timestep based on intermediate perturbed data. We formulate an optimization problem and derive an update rule that refines the text embeddings at each sampling step to improve alignment and preference between the mean predicted image and the text. This allows DATE to dynamically adapts the text conditions to the reverse-diffused images throughout diffusion sampling without requiring additional model training. Through theoretical analysis and empirical results, we show that DATE maintains the generative capability of the model while providing superior text-image alignment over fixed text embeddings across various tasks, including multi-concept generation and text-guided image editing. Our code is available at https://github.com/aailab-kaist/DATE.